cua_s1: complete native vision and screenshot inference with GPU parity - #64
Open
Levius-Fubuki wants to merge 12 commits into
Open
Levius-Fubuki wants to merge 12 commits into
Levius-Fubuki wants to merge 12 commits into
Conversation
# Conflicts: # recipe/cua_s1/native.md # src/models/cua_s1/native/Cargo.toml
…e-vision # Conflicts: # recipe/cua_s1/native.md # src/models/cua_s1/native/src/lib.rs
This was referenced Oct 2, 2026
This was referenced Oct 5, 2026
Levius-Fubuki
marked this pull request as ready for review
October 5, 2026 03:16
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Complete native screenshot inference: PNG/JPEG → RGB preprocessing → 24-block CUDA vision encoder and merger → image feature insertion and T/H/W positions → language execution → candidate decisions.
omni-cua-s1-visionserves the screenshot contract and reuses image features across questions in each request.Includes #59, #63 and the refreshed #56. Vision retains BF16 base weights and all 50 FP32 LoRA pairs separately; language uses a merged BF16 export. Startup verifies pinned checkpoint/export identity and hashes, then executes real warmup before readiness.
Refresh against current main's shared Qwen module, optimized kernels and 64-entry graph cache. CPU decoding/tokenization and response reconstruction live in the processor, outside executor admission. Shared
SerialScheduleradmits one complete request; both CUDA streams drain before release, including failures. Multimodal execution remains eager. CUDA ABI is 5: rebuild both library and worker.The diff retains runtime code, required dependencies, upstream licenses/notices and the one-time exporter. External tests, reports and controls are kept outside the PR diff.
Build and launch
Obtain the pinned weights/lock and Python environment using
recipe/cua_s1/text.md; additionally installtorchvision==0.29.0 numpy==2.5.3 Pillow==11.3.0.PYTHONPATH=src .venv/bin/python recipe/cua_s1/export_multimodal_language.py \ --base weights/Qwen3.5-4B --adapter weights/cua-s1-4b-0.2/multimodal \ --out weights/cua-s1-multimodal-language src/backends/cuda/qwen3_5/build.sh target/release 89 cargo build --release --locked -p omni-cua-s1-native --bins CUA_S1_BASE=weights/Qwen3.5-4B \ CUA_S1_VISION_ADAPTER=weights/cua-s1-4b-0.2/multimodal \ CUA_S1_MODEL=weights/cua-s1-multimodal-language \ CUA_S1_CUDA_LIB=$PWD/target/release/libqwen3_5_cuda.so \ target/release/omni-cua-s1-visionThe worker defaults to
127.0.0.1:8000;/healthreports readiness after initialization/warmup. Export manifest hashes supply local provenance, not an external signature; keep checkpoint files immutable during execution.Test Plan
Baseline:
f594d7dfc4c2bef812e23f7ed73573be9625b287.Head:
1152531a704044c879adabf377986d14010edbdb.Build source:
9538f00c5b3c8eb57b3bbb1c874f32c97b459070, whose tree is byte-identical to final head after merging #56 (f25e47a5bc082c3b94462bc67486f792bbfc3794). Remote production source hashes were verified against the final head.One RTX 4090 (24,564 MiB), driver 595.71.05, CUDA 13.0.88 / sm89, Rust 1.99.0, Python 3.12.3, PyTorch 2.14.0+cu130, Transformers 5.17.0, PEFT 0.21.0. Base
Qwen/Qwen3.5-4B@851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a; multimodal adaptercua-ai/cua-s1-4b-0.2@16818868b0cc7813808aae4e87b417657046ab79. All downloaded files passed pinned size/SHA-256 checks. Vision reference uses separate unmerged BF16/FP32 runs with TF32 disabled; native language uses the newly exported merged BF16 checkpoint.Frozen standard corpus: 7 requests / 8 questions, 1–26 candidates, PNG/JPEG, portraits/wide images and question reuse. Boundary corpus: 4 requests / 11 questions, including 1×1, 200×1, 1,024×1,024, noise and eight-question requests. Fresh BF16 and FP32 controls were regenerated on this host. Each fixed corpus must satisfy
max native probability error <= 2 * max BF16 reference error + 0.01, and match choices where FP32 margin ≥ 0.05. No sample or tolerance changes.Test Result
Local formatting, strict all-target Clippy, workspace tests and locked release build passed; current-head GitHub Rust, benchmark and docs CI passed (deployment skipped).
1e-7; nine invalid-input cases passed per corpus.Demo / evidence
These are fresh actual GPU/HTTP results for the verified final runtime tree, not the previous cleanup run. Raw output JSON, FP32/BF16 controls, frozen protocol, hashes, setup/build logs and reproduction scripts are retained in
artifacts/cua-native-refresh-20261005/and the matching execution-host directory. Current-head CI. External replay/control tools derive from the pre-cleanup source; updated graph/processor/concurrency harnesses remain local under the core-code constraint.This finite corpus establishes neither bitwise vision equivalence nor general model accuracy. Latency/throughput and cold-vs-warm performance were not benchmarked; no speed or memory improvement is claimed. Download retries and two PR56 harness/build-target failures are preserved in the evidence. Corrections changed only transfer/test setup, without runtime patches, discarded numerical failures or relaxed tolerances. No video is needed for this API parity run; raw assets are not publicly attached.
Self-review
Full diff and architecture contracts reviewed; independent source review found no actionable findings. Contributor checklist retained below.